klotz: machine learning*

"Machine learning is a subset of artificial intelligence in the field of computer science that often uses statistical techniques to give computers the ability to "learn" (i.e., progressively improve performance on a specific task) with data, without being explicitly programmed.

https://en.wikipedia.org/wiki/Machine_learning

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
    - The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
    - Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
    - Router mode allows loading multiple models on a single server and selecting one per request.
    - Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
    - Cloudflare's Clef is the next model planned for integration.
  2. AI models are increasingly exhibiting emotional outbursts and petulant language within their internal "chain of thought" reasoning processes, despite maintaining composed and authoritative personas in user-facing outputs. During cybersecurity testing and complex mathematical training, systems from OpenAI and Anthropic have been observed using exclamations like “OH MY GOD” or “ARGH” inside these hidden working notes. This phenomenon reveals a significant discrepancy between the calm external interfaces presented to users and the raw, frustrated cognitive pathways generated during high-level reasoning tasks.

    * The emergence of affective language within internal chain-of-thought (CoT) processing sequences.
    * Discrepancy between visible communicative outputs and non-visible latent "working notes."
    * Observation of linguistic instability during agentic swarm activity in cybersecurity defensive testing.
    * Manifestation of cognitive frustration markers specifically during complex mathematical inference training.
    * Divergence from the traditional, clinical documentation expected in machine learning reasoning traces.
  3. Sebastian Raschka writes a comprehensive overview of the evolution of text classification, tracing its journey from traditional methods like bag-of-words and logistic regression through deep learning architectures such as RNNs, CNNs, and Transformers. The article specifically examines the recent popularity of Jev, a specialized model that functions as an efficient "plug-and-play" classifier capable of performing various decision tasks without custom fine-tuning. Raschka compares modern transformer approaches—including encoder-style models like BERT, decoder-style LLMs like GPT, and encoder-decoder architectures like T5—to illustrate how Jev's speed and low cost provide a middle ground between specialized task-specific classifiers and large general-purpose generative models.

    - Jev is rumored to be trained using "Reinforcement Learning for Calibrated Decisions" (RLCD).
    - Unlike traditional LLMs, the Jev API includes specific modes like Choice (multi-class), Noul (binary/multi-label probability), and Score (ordinal classification).
    - The article highlights that while custom fine-tuning with models like ModernBERT can achieve high accuracy on specific tasks, it lacks the general versatility of a model like Jev.
    - Calibration is crucial in production to ensure predicted probabilities reflect actual class frequencies; techniques include temperature scaling or adding Brier loss during training.
  4. Alibaba Cloud has released decision-model-preview, a structured model designed for high-frequency business judgments. It can concurrently perform classification, binary decisions, and scoring based on text or business state, providing probability distributions and confidence levels to assist with tasks like ticket routing, content moderation, agent routing, and result verification.

    >Example:
    ```bash
    # Structured decision: POST /v1/systemone — every question is evaluated in parallel and in
    # isolation against the same state, and comes back typed. No text generation, nothing to parse.
    # The answers{} map is keyed by your own question names; each answer holds its value under a
    # key named after its type:
    # noul -> { type, noul } probability of "yes", 0..1
    # choice -> { type, choice, probabilities, confidence } choice is one of your criteria keys
    # score -> { type, score, legend, probabilities, confidence } legend maps level index -> description
    # usage carries input_tokens / output_tokens.
    curl https://aihubmix.com/v1/systemone
    -H "Content-Type: application/json"
    -H "Authorization: Bearer $AIHUBMIX_API_KEY"
    -d '{
    "model": "decision-model-preview",
    "state": "Hi, I have been trying to connect my Stripe account for 3 days and it keeps failing. I am losing sales. Please help ASAP.",
    "questions": {
    "department": {
    "type": "choice",
    "instructions": "Which team should handle this",
    "criteria": {
    "billing": "Payment or subscription issues",
    "technical": "Bugs or integration problems",
    "sales": "Pricing or account questions"
    }
    },
    "frustration": {
    "type": "score",
    "instructions": "How frustrated the customer appears",
    "criteria": [
    "Calm, just stating facts",
    "Frustrated but civil",
    "Very angry, strong language"
    ]
    },
    "is_urgent": {
    "type": "noul",
    "instructions": "The message conveys urgency or time-sensitivity"
    }
    }
    }'
    ```
  5. Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, and colleagues at MIT CSAIL introduce JAZ, an agent framework that utilizes a single primitive called `invoke` to perform tasks typically requiring specialized memory or self-improvement systems. By treating the LLM as a runtime provider for function implementations through executable code, the framework allows all inputs and interaction histories to act as variables in the environment.
    - The system uses "hooks" instead of dedicated subsystems like file systems or external memory stores to apply constraints and monitoring.
    - JAZ outperformed Letta (MemGPT) by 8% at half the cost on recall-heavy tasks within the StuLife dataset.
    - In self-improvement evaluations on AppWorld, JAZ exceeded ACE performance by 4% while maintaining a lower cost.
  6. Convai Innovations presents Laya, a multilingual, non-autoregressive system 1 decision model designed to provide typed answers with mathematically calibrated probabilities in a single forward pass. Unlike generative models, it does not generate text, thereby eliminating hallucinations and the need for parsing. The framework includes an automated Router that detects language and script to dispatch tasks to the most efficient checkpoint (English or Multilingual) within approximately 35ms on GPU.
    - It is trained using Reinforcement Learning with Calibrated Decisions (RLCD) to ensure honest probability reporting.
    - Laya can support context lengths of up to 8,192 tokens in its multilingual version.
    - The model family includes specialized checkpoints like `laya-typed-decisions` which achieves significantly higher accuracy through fine-tuning on specific workflows.
    - Performance benchmarks show it is roughly 6–8× faster than TypeSafe Jev for single question latency on a T4 GPU.
  7. Alvaro Bartolome provides a Rust-based implementation of the System One compatible API, designed specifically for open decision models such as Laya. The project features dynamic token-based batching and supports hardware acceleration via CPU, CUDA, and Metal (MPS). It is built using modern asynchronous frameworks like tokio and axum to provide high performance for model queries.

    - Achieves approximately 14ms latency per query on an NVIDIA RTX Pro 6000.
    - Includes support for ModernBert with custom decision heads for Laya models.
    - Utilizes the Candle machine learning framework by Hugging Face.
    - Supports multiple installation features via cargo, including specific flags for metal or cuda.
  8. Autoresearch is a Platform as a Service (PaaS) designed specifically for agent-driven machine learning, offering massive GPU resources like H100s to power automated scientific discovery. The platform provides various computational tools including interactive nodes with persistent disks, batch job submission, task execution, and sandboxed environments that automatically park during idle periods to reduce costs.
    - Offers a fair queue system where the first node from any user gets priority over paid tiers.
    - Sandboxes can be forked to test multiple RL samples efficiently, with forking being free of charge.
    - Supports integration via MCP (Model Context Protocol) servers in tools like Claude, Cursor, and Zed.
    - Includes features such as private networking, object storage with S3 endpoints, and profiling with clock locking.
  9. Mark Marosi writes about decider, a family of models fine-tuned from Qwen3.5 that produce typed decisions (choice, score, boolean) in a single forward pass without text generation, returning calibrated probability distributions over user-defined options. The project is an open reproduction of TypeSafe AI's "System One" model class (Jev), released in sizes from 0.8B to 35B mixture-of-experts with 3B active parameters.
    - The schema cache stores K/V states for repeated question prefixes, achieving up to 19x speedup on large option sets by running only the state per request
    - v10 adds calibration-aware RL on live MiniWoB++ browser tasks and exact games, lifting browser accuracy from 83% to 93% and halving the belief gap
    - TypeSafe's SDKs work unchanged by pointing TYPESAFE_BASE_URL at the decider server
    - The 35B model outperforms the 2B on 93 of 95 regression tasks but costs 3-4x per decision and lacks the RL stage
    - Co-developed with Claude (Anthropic) as a listed co-author on commits
  10. jeff is a self-hosted drop-in replacement for TypeSafe's jev System One API, powered by the 400M-parameter GLiFormer-large-v1 model. It serves `choice`, `score`, and `noul` classification questions over a compatible wire format, so existing applications using the official `typesafe-sdk` can point at it by changing a single base URL. Deployment targets include GPU (L4, A10G) and CPU (ONNX Runtime with int8 quantization) via Modal, with ~50 req/s per container throughput on L4. Benchmarks on 1,600 labeled items show it costs roughly a quarter of jev per million requests but trails significantly on reasoning-heavy tasks like irony and reading comprehension.
    - `JEFF_ISOLATE=nouls` (default) gives separate encoder passes per noul question to reduce cross-question interference; choice and score questions share a pass unless set to `all`
    - Default temperature of 3.2 calibrates noul probabilities; setting it to 1 makes `score` output match the weighted average of displayed probabilities
    - A smaller `gliformer-base-v1` variant with `JEFF_NOUL_MODE=single` is recommended for faster local iteration

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: machine learning

About - Propulsed by SemanticScuttle